前三天我們有了讀、分塊、呼叫的能力。但系統還缺一個致命的東西:記憶。
想象一份 100 頁的規格書。第 2 頁寫著:
REQ-2.1.1: 系統支持多用戶並行存取
翻到第 90 頁:
REQ-8.4.2: 系統採用單用戶模式,每次只有一個用戶活動
人工審查時,你會立刻看到矛盾。但 LLM 呢?
如果我們按 Day 3 的分塊邏輯處理,LLM 分析第 2 章時會看到 REQ-2.1.1。分析完後,結果被丟棄。分析第 8 章時,LLM 看到 REQ-8.4.2,但根本想不起來 REQ-2.1.1。
結果:矛盾被漏過了。
這不是 LLM 笨。這是我們的架構設計沒有給它「記住」的機制。
核心思想很簡單:不要每次分析都丟棄結果,而是把所有約束存在一個不斷增長的池裡。
初始化
↓
[全局約束池(空)]
讀第 1 個 chunk(第 2 章)
提取:REQ-2.1.1, REQ-2.1.2, ...
添加到池
池 = {REQ-2.1.1, REQ-2.1.2}
↓
讀第 2 個 chunk(第 3 章)
提取:REQ-3.2.1, ...
對照池 → 無衝突
添加到池
池 = {REQ-2.1.1, REQ-2.1.2, REQ-3.2.1}
↓
... 繼續 ...
↓
讀第 N 個 chunk(第 8 章)
提取:REQ-8.4.2
對照池 ✅ 發現!REQ-8.4.2「單用戶」vs REQ-2.1.1「多用戶」
記錄矛盾
添加到池
↓
生成報告:3 個約束,1 個矛盾
關鍵:約束池永遠在那裡。每個新約束都能看到之前的所有約束。
三個核心概念:
Constraint(約束)— SRS 裡的一個需求:
@dataclass
class Constraint:
id: str # REQ-2.1.1
text: str # "系統支持多用戶並行存取"
req_type: str # "功能需求"
chapter: int # 2
chunk_id: int # 0
dependencies: List[str] # 依賴其他需求的 ID
conflicting: List[str] # 已知衝突的 ID
line_number: int # 行號
Conflict(矛盾)— 兩個約束之間的衝突:
@dataclass
class Conflict:
req_id_1: str # "REQ-2.1.1"
req_id_2: str # "REQ-8.4.2"
conflict_type: str # "邏輯矛盾"
description: str # 具體描述
severity: str # "高", "中", "低"
location_1: str # "第 2 章"
location_2: str # "第 8 章"
suggested_fix: str # 修復建議
LangGraphState(全局狀態)— Agent 的整個狀態機:
class LangGraphState(TypedDict):
# 輸入
srs_file_path: str
srs_content: str
# 處理
chunks: List
current_chunk_id: int
# ⭐ 核心:約束池
global_constraints: List[Constraint]
# 結果
conflicts: List[Conflict]
analysis_trace: List[str] # 分析日誌
# 狀態
processing_status: str # "初始化中" / "處理中" / "完成"
error_message: Optional[str]
src/models.py:
from dataclasses import dataclass, field
from typing import List, Optional, TypedDict
from enum import Enum
class ReqType(str, Enum):
FUNCTIONAL = "功能需求"
NON_FUNCTIONAL = "非功能需求"
SECURITY = "安全需求"
PERFORMANCE = "性能需求"
class ConflictType(str, Enum):
LOGIC = "邏輯矛盾"
PERFORMANCE = "性能衝突"
DEPENDENCY = "依賴缺失"
AMBIGUITY = "模糊性"
INCOMPATIBILITY = "不兼容性"
class Severity(str, Enum):
HIGH = "高"
MEDIUM = "中"
LOW = "低"
@dataclass
class Constraint:
"""需求規格書中的一個約束"""
id: str
text: str
req_type: ReqType
chapter: int
chunk_id: int
dependencies: List[str] = field(default_factory=list)
conflicting: List[str] = field(default_factory=list)
line_number: int = 0
def __str__(self) -> str:
return f"{self.id}: {self.text[:50]}..."
@dataclass
class Conflict:
"""需求之間的矛盾"""
req_id_1: str
req_id_2: str
conflict_type: ConflictType
description: str
severity: Severity
location_1: str = ""
location_2: str = ""
suggested_fix: Optional[str] = None
def __str__(self) -> str:
return (
f"{self.severity.value} - {self.conflict_type.value}: "
f"{self.req_id_1} vs {self.req_id_2}"
)
@dataclass
class AnalysisResult:
"""分析結果"""
srs_filename: str
total_constraints: int
conflicts: List[Conflict]
analysis_trace: List[str]
def summary(self) -> str:
high = len([c for c in self.conflicts if c.severity == Severity.HIGH])
medium = len([c for c in self.conflicts if c.severity == Severity.MEDIUM])
low = len([c for c in self.conflicts if c.severity == Severity.LOW])
return (
f"分析完成:{self.total_constraints} 個約束,"
f"{len(self.conflicts)} 個矛盾 (高:{high}, 中:{medium}, 低:{low})"
)
class LangGraphState(TypedDict):
"""Agent 的全局狀態機"""
srs_file_path: str
srs_content: str
chunks: List
current_chunk_id: int
global_constraints: List[Constraint]
conflicts: List[Conflict]
analysis_trace: List[str]
processing_status: str
error_message: Optional[str]
想象分析一份規格書的流程:
def analyze_with_state(agent: OllamaAgent, srs_text: str) -> LangGraphState:
"""用 Global State 分析 SRS"""
# 1. 初始化 State
state: LangGraphState = {
"srs_file_path": "srs.md",
"srs_content": srs_text,
"chunks": agent.chunk_srs(srs_text),
"current_chunk_id": 0,
"global_constraints": [], # ⭐ 空的約束池
"conflicts": [],
"analysis_trace": ["[初始化]"],
"processing_status": "處理中",
"error_message": None,
}
# 2. 遍歷每個 chunk
for chunk_id, chunk in enumerate(state["chunks"]):
state["current_chunk_id"] = chunk_id
# 3. 提取這個 chunk 裡的約束
constraints = extract_constraints_from_chunk(chunk, agent)
# 4. ⭐ 關鍵:對照全局約束池
for new_constraint in constraints:
for existing in state["global_constraints"]:
if conflicts_with(existing, new_constraint):
# 發現矛盾!
conflict = Conflict(
req_id_1=existing.id,
req_id_2=new_constraint.id,
conflict_type=ConflictType.LOGIC,
description=f"{existing.text} vs {new_constraint.text}",
severity=Severity.HIGH,
location_1=f"第 {existing.chapter} 章",
location_2=f"第 {new_constraint.chapter} 章",
)
state["conflicts"].append(conflict)
# 5. 把新約束加入全局池
state["global_constraints"].extend(constraints)
# 6. 記錄進度
state["analysis_trace"].append(
f"[Chunk {chunk_id}] 提取 {len(constraints)} 個約束,"
f"發現 {len(state['conflicts'])} 個矛盾"
)
state["processing_status"] = "完成"
return state
每次 loop,狀態都更新一次。約束池會越來越大,矛盾越來越多被發現。
簡單說:沒有它,系統檢不出跨章節的矛盾。
有了它:
Day 5 我們用 LangGraph 的 StateGraph 把這個手工編寫的邏輯變成真正的狀態機。現在是手動流轉,Day 5 之後會自動流轉。
第 1 週:透明的 SRS 分析
約束池的思路對,但每新增一條就跟整池 pairwise 比對,規格書一長 LLM 呼叫次數會長平方級,而且大多數比對的都是無關條文。我會把每條約束先正規化成結構化欄位(主體/屬性/取值),再用 embedding 只撈語意最近的 K 條出來給 LLM 判斷衝突,候選少了判斷也會更準。提取時順便記下頁碼和 REQ 編號,最後的矛盾報告才拿得出證據回查原文。